Journal of Proteome Research
● American Chemical Society (ACS)
Preprints posted in the last 90 days, ranked by how well they match Journal of Proteome Research's content profile, based on 234 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit.
Schiebenhoefer, H.; Muth, T.; Fuchs, S.; Renard, B. Y.
Show abstract
Metaproteomics is the investigation of the protein composition of multi-organism samples. While metagenomics answers the question which organisms are present in a sample, metaproteomics additionally answers the question which organisms are active. State-of-the-art tools for annotating proteomic data with taxonomic information (e.g. Unipept, DIAMOND) do not control the false taxonomic identification rate, which can lead to incorrect results and thus incorrect interpretations, as we demonstrate with examples. ProteoDUDes processes the results from popular sequence annotation tools so that the proportion of true identifications in the result is at as high as or higher than in the the compared tools. We evaluate ProteoDUDes on simulated data and experimental mock community data. Our results indicate that ProteoDUDes has the same error rate as other tools on simulated data and half the error rate on the experimental mock community data. This allows more accurate statements to be made about which organisms are functionally active in a complex sample. ProteoDUDes is open-source and available at https://github.com/pirovc/dudes.
Juber, M.; Binti, S.; Pandi, B.; Lau, E.; Lam, M. P. Y.
Show abstract
Alternative splicing is an important regulatory layer in gene expression, but knowledge on the isoform protein molecules continue to lag their canonical counterparts. An open question is whether alternative protein isoforms feature different half-life than the canonical counterpart, which could indicate differential usage and functional diversification. Here we combined a proteogenomics approach with heavy water-based protein turnover analysis to survey 24 pairs of canonical-alternative protein isoforms in the mouse heart. The results provide a reference on their numerical half-life and also reveal widespread differences in isoform stability.
Mayer, R. L.; Mechtler, K.
Show abstract
While the field of immunopeptidomics has matured substantially over the last years, high input amounts of cellular or tissue material are still required to obtain a somewhat complete profile of the immunopeptidome. Here we present a simple platform termed TripleToolWF (derived from Triple Tool workflow) to increase the number of identified and quantified immunopeptides combining the outputs of three search engines such as PEAKS Online 12, Sequest HT with INFERYS rescoring and MSFragger. For assessing the false discovery rate (FDR) an entrapment approach is used. The platform improved peptide identifications by 6-14% and peptide quantitations by 11-25% compared to the best individual search engine for two independent, previously published, bacterial infection datasets. Peptides were mostly 9-12mers as expected and >90% of the obtained 9mers were predicted binders by the stringent majority voting approach of Immunolyser 2.0 which indicates high confidence of the identified immunopeptides. The FDR was monitored using dedicated entrapment searches against shuffled databases. The resulting entrapment FDR was assessed before and after result pooling and showed only a minor increase upon pooling compared to the worst individual search engine. It remained even below the target of 1% peptide FDR in 40% of the experiments. Compared to the original publications, the number of high confidence bacterial immunopeptides was drastically elevated by 53% and 2800% for the Listeria monocytogenes and Mycobacterium bovis BCG projects, respectively, when applying strict filters. Of these additional bacterial sequences, all 9mer sequences were predicted as binders by at least one of the prediction algorithms of Immunolyser 2.0 illustrating their actual HLA binding nature. TripleToolWF hence provides a simple tool to further increase the number of obtained sequences from MS-based immunopeptidomics experiments to facilitate a deeper view of the immunopeptidome for refined vaccine candidate prioritization.
Kotli, P. C.
Show abstract
Ancient DNA (aDNA) has transformed the study of hominin relationships, but its preservation in ancient fossils is often limited. Enamel palaeoproteomics offers an alternative molecular approach for taxonomic analysis. In this study, we re-analyse published DDA mass spectrometry data1 from the Denisovan-attributed Penghu 1 mandible (PXD054412)2 and a Neandertal enamel specimen from Gruta de Oliveira, Portugal (PXD038154) 3. Five AMBN peptides carrying the Valine-273 substitution (V273) were validated in the Penghu 1 enamel, three of which were independently detected in both DDA acquisitions. No V273-containing peptide signal was detected in the Neandertal dataset. Conversely, the ancestral Methionine-273 peptide REDPM[+16]AYG was detected exclusively in the Neandertal specimen. Extracted ion chromatograms, isotopic envelope confirmation, and MS2 fragmentation spectra, all support the reported peptide assignments. AMELY-specific peptides additionally support male sex assignment for both ancient individuals. Together, these results confirm the taxon-specific mutual exclusivity of AMBN V273 and M273 variants across Denisovan and Neandertal lineages, establishing AMBN M273V as a molecularly validated diagnostic marker for Denisovan identification from dental enamel. More broadly, targeted MS1 reanalysis of public proteomics datasets provides a scalable complement to ancient genomics for resolving hominin lineage identity and sex determination when DNA is not preserved
Nameni, A.; Declercq, A.; Gabriels, R.; Degroeve, S.; Martens, L.; Bouwmeester, R.
Show abstract
In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly supports this task through peptide-spectrum match (PSM) rescoring, in which a classifier, typically a linear semi-supervised model, refines the initial matching score. However, Mokapot allows the user to choose among different machine learning algorithms of increasing complexity, from the default linear support vector machine (LSVM) to random forest and XGBoost. Here, we use an entrapment approach to assess the effect of this increasing complexity on PSM identification and the accuracy of the estimated false discovery rate (FDR). We show that, while more complex models increase the number of identified PSMs at a fixed FDR threshold, this gain reflects a bias towards random matches from the target proteome database rather than genuine identifications. Indeed, for the most complex model, the entrapment FDR reaches 6.3% instead of the estimated 1% decoy FDR. This bias thus yields overly optimistic FDR estimates, indicating that model complexity in PSM rescoring must be carefully balanced against this overfitting risk.
Wen, B.; Li, K.; Riffle, M.; MacCoss, M. J.; Bittremieux, W.; Noble, W. S.
Show abstract
De novo peptide sequencing detects peptides directly from tandem mass spectra without a protein sequence database, and deep learning has substantially advanced its performance. Casanovo, one such widely used model, is distributed as a Python command-line program. Consequently, installation, GPU and dependency configuration, and manual parameterization can be challenging for many bench scientists and are a recurring source of errors. Interpreting and validating the resulting predictions poses a further challenge. We present CasanovoGUI, an open-source Java-based desktop application that makes all of Casanovos main analysis functions available through a point-and-click interface on Windows, macOS, and Linux. On first use, CasanovoGUI automatically installs a private Python environment and Casanovo with a GPU-matched build, requiring no prior software setup. The GUI provides access to Casanovos analysis functions and configuration parameters, streams live progress, and integrates results interpretation: annotated spectra with per-residue confidence scores in the PDV viewer, and mismatch-tolerant mapping of de novo peptides back to a reference proteome. CasanovoGUI is available at https://github.com/Noble-Lab/CasanovoGUI.
Fields, L.; Hubecky, E. M.; Selby, K. G.; Li, L.
Show abstract
Data-independent acquisition (DIA) mass spectrometry has emerged as a powerful tool for neuropeptidomics, but its success relies heavily on the quality of spectral libraries used for peptide identification. There are inherent challenges to mass spectrometry analysis of crustacean neuropeptides, including the endogenous nature in which they are analyzed, extensive post-translational modification (PTM), and atypical fragmentation patterns. Thus, general-purpose proteomic spectral prediction tools may not perform optimally in the endogenous peptide domain. In this study, we benchmark four widely used spectral prediction platforms, Prosit, MS2PIP, AlphaPeptDeep, and UniSpec, to evaluate their performance in predicting the fragmentation of neuropeptides. Using an empirically derived spectral library from crustacean tissues as reference, we assess model compatibility, dot-product similarity, Pearson correlation, and DIA-based identifications across brain, sinus gland, and pericardial organ samples. Our results reveal that no single model comprehensively captures neuropeptide fragmentation characteristics. While UniSpec showed unexpected strengths due to its inclusion of neutral loss ions, AlphaPeptDeep demonstrated the highest spectral similarity, and MS2PIP and Prosit outperformed in DIA-NN identifications. We further highlight the critical impact of neutral loss fragments, present in over 50% of empirical spectra, and emphasize the need for hybrid spectral libraries that integrate complementary strengths across models. This work provides a foundational framework for optimizing spectral library selection in neuropeptidomics and underscores the importance of model-specific biases when analyzing structurally diverse endogenous peptides.
Vasylieva, V.; Massignani, E.; Claeys, T.; Bourassa, F.; Leblanc, S.; Arefiev, I.; Martens, L.; Brunet, M. A.
Show abstract
ShortThe SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. LongO_ST_ABSBackgroundC_ST_ABSThe SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. ResultsCompared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. ConclusionsIn this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.
Yue, Q.-X.; Wei, Z.; Dai, C.; Bai, M.; Perez-Riverol, Y.; Sachsenberg, T.
Show abstract
With the rapid development of mass spectrometry-based proteomics, the volume of phosphoproteomic data has increased substantially. However, accurate localization of phosphorylation sites and standardized statistical validation remain critical analytical bottlenecks. To address the lack of standardized cross-algorithm evaluation, we introduce onsite, a unified and open-source Python framework. onsite integrates an alanine-decoy strategy to estimate the false localization rate (FLR) across three algorithms: AScore, PhosphoRS, and pyLucXor. This modular architecture efficiently processes large-scale datasets and enables global FLR calculation. Benchmarking on the standard synthetic phosphopeptide dataset PXD000138 highlighted distinct inter-algorithmic variations. Using the same 5% global FLR threshold, pyLucXor localized the most target sites (28,353). It also reached a high accuracy (91.22%) against the known ground truth, resulting in the largest number of correctly localized sites (25,865). Reanalysis of the highly fractionated, large-scale PXD012255 dataset further demonstrated that native integration of onsite into the quantms pipeline enables scalable processing and provides a standardized framework for FLR control in large-scale phosphoproteomics. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=64 SRC="FIGDIR/small/737157v1_ufig1.gif" ALT="Figure 1"> View larger version (14K): org.highwire.dtl.DTLVardef@e4c85dorg.highwire.dtl.DTLVardef@1e8464org.highwire.dtl.DTLVardef@185cea1org.highwire.dtl.DTLVardef@1c0d1bc_HPS_FORMAT_FIGEXP M_FIG C_FIG
de Almeida, R. F.; Fernandes, M.; de Godoy, L. M. F.
Show abstract
The processes such as DNA replication, transcription, and repair are often modulated by specific nuclear proteins, protein-protein interactions (PPIs), and post-translational modifications (PTMs). In Trypanosoma cruzi, however, the nuclear proteome and interactome have not been systematically mapped, limiting the interpretation of nuclear regulatory processes. Here, we report a nuclear proteome resource generated from intact nuclei isolated from T. cruzi and analyzed by high-resolution Orbitrap LC-MS/MS, integrating proteome profiling, computational interaction network inference, and exploratory crosslinking mass spectrometry (XL-MS). Proteome profiling identified 1,734 proteins in the nuclear fraction, including 316 proteins identified with PTM-containing peptides. Subcellular localization prediction and Gene Ontology analysis support nuclear enrichment and highlight functions related to transcription, RNA metabolism, and genome maintenance. The in silico interaction network derived from STRINGDB organizes the proteins into functional clusters, including a histone-associated interaction neighborhood. In parallel, XL-MS identified 26 residue-resolved interprotein crosslinks involving 36 proteins and detected PTMs at or near linked residues. Together, these data support reuse for comparative nuclear proteomics, multi-omics integration, and prioritization of candidates for future functional studies.
Brenes, A. J.; Mayer, R. L.; Makar, A.; Coelho, P.; van Stralen, G.; Sadiku, P.; Walmsley, S. R.; Matzinger, M.; Mechtler, K.; von Kriegsheim, A.
Show abstract
Mass spectrometry-based single cell proteomics (SCP) is rapidly emerging as a powerful approach for biological research, with applications extending beyond in-vitro cancer cell lines. Recent advances make it possible to apply SCP to ex-vivo human cells from tissues such as the brain and pancreas, as well as to technically challenging immune populations such as neutrophils. However, these analyses remain more challenging and typically result in reduced proteomic coverage. To support the development of robust workflows for SCP data acquisition and analysis, we systematically evaluated multiple DIA search engines, search engine settings, the inclusion of high-load library samples in single-cell search spaces, the impact of contaminants, and the quantitative properties of identified proteins. These comparisons were performed across two major instrumentation platforms, Orbitrap Astral and timsTOF SCP, and across A549, RKO cells and neutrophils, three cell types differing in size and protein content. Our work here provides guidelines on the software parameters to use for SCP, instrument specific results and cell dependent optimizations of high-load libraries, as well as novel evaluation of the quantitative properties of proteins for single cell and low input proteomics.
Kotimoole, C. N.; Arefian, M.; McKay, E. C.; Kasaragod, S.; Skoraczynski, G.; Collins, B. C.
Show abstract
SummaryContemporary proteomics methods can now generate large-scale DIA datasets of thousands of files that demand substantial computational resources for efficient analysis. Spectronaut is a widely used platform for DIA data processing; however, large-scale searches are often constrained by computational performance and long execution times when run on single workstations. Here, we present Spectronaut-nf, a Nextflow-based pipeline that enables scalable and parallelized execution of Spectronaut analyses across high-performance computing (HPC) environments. The workflow divides directDIA analysis into modular stages, including spectral library generation, DIA searching, and merging results, allowing efficient distribution of tasks across multiple compute nodes. Benchmarking using 72 diaPASEF raw files using typical hardware demonstrated that Spectronaut-nf completed searches in 23.77 hours, compared with 39.09 hours on a Windows workstation and 67.04 hours on a single-node Linux HPC setup. Stress testing with 1,037 diaPASEF raw files further demonstrated the scalability and robustness of the workflow for large proteomics datasets. Across platforms, protein and peptide identifications remained consistent, with only minimal variability attributable to platform-specific differences. Overall, Spectronaut-nf provides a flexible, scalable, and efficient framework for high-throughput DIA proteomics analysis in HPC environments. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=129 SRC="FIGDIR/small/741433v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@4a7107org.highwire.dtl.DTLVardef@14287d8org.highwire.dtl.DTLVardef@e484e9org.highwire.dtl.DTLVardef@d21c07_HPS_FORMAT_FIGEXP M_FIG C_FIG
Mathai, D.; Schulze, S.
Show abstract
Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze datasets on the scale of whole proteomes. Benchmarking on a curated dataset of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant datasets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics dataset of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic datasets and whole proteomes. For Table of Contents Only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=98 SRC="FIGDIR/small/737882v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@1011179org.highwire.dtl.DTLVardef@1222493org.highwire.dtl.DTLVardef@8f69f2org.highwire.dtl.DTLVardef@174b30e_HPS_FORMAT_FIGEXP M_FIG C_FIG
Feltenstein, I. G.; Drown, B. S.
Show abstract
Proteins are dynamically regulated by a myriad of post-translational modifications (PTMs) that control their stability, conformation, activity, subcellular localization, and local interactions. Capturing the precise composition of these various modification states, or proteoforms, is a principal objective of top-down proteomics (TDP). By ionizing intact proteoforms and combining measurements of precursor ion and fragment ion masses, the position, stoichiometry, and combination of PTMs can be determined. Despite the highly valuable measurements that TDP can provide, it is typically less sensitive than corresponding peptide-level analysis with many reports utilizing input material in the microgram to milligram range. Contributing to this lack of sensitivity is the risk of sample loss due to non-specific binding to surfaces during sample preparation. The most widely employed sample preparation approaches for TDP either require high sample input (e.g. precipitation and ultra-filtration) or fail to effectively remove surfactants (e.g. solid-phase extraction). These limitations have hindered advancement of targeted TDP applications involving immunoprecipitation and other enrichment strategies. Bead-assisted protein aggregation, also referred to as single-pot, solid-phase-enhanced sample preparation (SP3), has emerged as a popular sample preparation strategy for bottom-up proteomic workflows, but has only been used in TDP with secondary ion exchange chromatography cleanup. We envisioned a magnetic bead based protein cleanup approach that proceeds directly to MS analysis with judicious choice of bead surface chemistry and elution conditions. Here we report a sample preparation method using hydroxyl-functionalized magnetic beads for top-down proteomics applications.
Wu, A.; Kohler, D.; Navada, P. P.; Robbins, J. E.; Boyle, G. E.; Boshart, A.; Karis, K.; Neefjes, J.; Konvalinka, A.; Sarthy, J.; Pino, L.; Gyori, B. M.; Vitek, O.
Show abstract
A common outcome of quantitative mass spectrometry-based proteomic and phosphoproteomic experiments is a list of proteins that are differentially abundant between conditions. However, biological interpretation requires evaluation in the context of prior knowledge of biological mechanisms and protein function. One approach to facilitate mechanistic biological interpretation is to integrate such lists with biological network databases, built from manually curated resources and text mining systems. This manuscript automates this process with MSstatsBioNet, a Bioconductor package that integrates MSstats, a family of open-source packages for detecting differentially abundant proteins, and INDRA, a system that extracts biomolecular networks from biomedical literature using text mining and merges those networks with the content of curated knowledge bases. Taking as input a list of differentially abundant proteins from MSstats, MSstatsBioNet retrieves a protein subnetwork from INDRA and overlays experimental fold changes onto the underlying subnetwork. Users can then interact with the network and overlaid data, interrogating primary literature evidence to construct granular mechanistic narratives for iterative hypothesis generation. We demonstrate the utility of this approach with three case studies, two measuring changes in protein abundance and one measuring changes in phosphorylation.
Hartmaring, Y.; Wang, S.; Jones, A. R.; Vizcaino, J. A.; Schlaffner, C. N.; Renard, B. Y.
Show abstract
Dysregulation of post-translational modifications (PTMs) is associated with severe pathologies, including cancers and Alzheimers disease. Despite their biological importance, identifying modified peptides remains challenging due to the immense combinatorial search space. While searches benefit from prior knowledge of a peptides modification status, the data scarcity for most PTMs hinders the development of accurate deep learning classifiers like AHLF (ad hoc learning of peptide fragmentation). Here, we overcome this data bottle-neck for acetylation and ubiquitination. We harmonised a dataset with about 500,000 high quality acetylated peptide-spectrum matches (PSMs) from nine publicly available acetylation-enriched datasets. We fine-tuned AHLF with the acetylation and a 2-million spectra strong ubiquitination dataset separately and assessed the minimum data requirement for training by iteratively downsampling. Training separate models on SILAC and label-free subsets also assessed the impact of data diversity. The resulting acetylation and ubiquitination models achieve an AUC of 0.87 and 0.90 respectively. Beyond 28,500 acetylated spectra, corresponding to roughly 0.3% of the original models training data, additional data just provides minor performance gains. Finally, we show that data diversity is beneficial for generalizability, while models trained on homogeneous data sources tend to overfit to their respective data type. All code, and model weights are available at https://gitlab.com/dacs-hpi/ahlf-ptmai.
Kalaidopoulou Nteak, S.; Drouin, N.; Michalik, S.; Salazar, M. G.; Hammer, E.; Dhople, V.; Holtfreter, S.; Weiss, S.; van den Toorn, H.; Broeker, B.; Domanska, G.; Voelker, U.; Heck, A.
Show abstract
The plasma proteome is a valuable resource for assessment of the physiological state of the donor. Containing hundreds of different proteins of variable concentrations, it displays substantial inter-donor differences in individual protein levels, making each plasma proteome highly donor-specific. Less is known about intra-donor variability in the plasma proteome over time, although such variations may even be more indicative of a changing physiological state. Here we assessed data obtained from the TIMES cohort, comprising 51 apparently healthy participants monitored monthly over 12 months, focusing especially on temporal variations in blood protein levels. Most strikingly, we observed that several women in this cohort revealed strongly correlated temporal variations in their plasma proteome, including most notably PZP, SHBG, FETUB, AGT, SERPINA6, SERPINA7, CP, APOL1 and KNG1, with levels sometimes fluctuating by more than 20-fold. In contrast, such variations were absent in men. Some of the fluctuating proteins have been known to be hormone-regulated (e.g., PZP, SHBG), but for others this was not yet fully clear. Through the tight co-variation observed for these proteins in the plasma proteome of women, we can conclude that all these proteins are similarly hormone regulated. The findings reported here not only corroborate previous studies showing estrogen-dependent regulation of several plasma proteins, but also extend this category to include also CP, APOL1, and KNG1. As these latter have been often proposed as candidate biomarkers, they should be validated in sex-balanced cohorts and interpreted with caution, especially in large-scale plasma proteomics studies wherein often only one or a few sampling time points are measured per donor.
Haueis, J. R. S.; Lazar, I. M.
Show abstract
Mass spectrometry (MS) is the leading technology for identifying proteins in complex biological samples. It relies on the use of tandem MS alongside a reference database of canonical protein sequences to computationally identify peptides and their parent proteins. The canonical sequences represent the most widely expressed and functionally validated forms of proteins. Consequently, disease-induced or disease-supportive variants, such as those associated with cancer, will evade detection if they are absent from the database. To address this challenge, this study introduces a revised release of the Unkown Mutation Analysis (XMAn) database by incorporating coding missense and nonsense mutations from the latest versions (v103) of the COSMIC Genome Screen Mutants (GSM) and Cancer Gene Census (CGC) datasets in two distinct FASTA-formatted peptide databases comprising 3,848,499 and 312,658 variants, respectively. The mutated peptides were matched to reviewed, non-redundant UniProt Homo sapiens protein entries (18,362 and 746), and characterized in terms of nucleotide- and amino acid mutation frequencies, peptide length distributions, and associations between specific single-nucleotide (SNV) and single amino acid (SAAVs) variants. Applied to the analysis of MDA-MB-231 breast cancer cell-membrane protein fractions, the database enabled the identification of 300+ high-quality variant peptides - several localized to functional protein-binding and catalytic domains - and 23 aberrant protein products mapped to the CGC dataset. The database is hosted and available for download on Zenodo (XMAn/gsm doi: 10.5281/zenodo.21781023; XMAn/cgc doi: 10.5281/zenodo.21781514) or can be accessed through https://sites.google.com/vt.edu/xman-db/home.
Benacom, D.; Specht, A.; Nicholas, J. C.; Guillard, R.; Gillman, M.; Dubin, R.; Ganz, P.; Rotter, J. I.; Taylor, K. D.; Rich, S. S.; Liu, P. Y.; Wood, A. C.; Mi, M. Y.; Deo, R.; Zitting, K.-M.; Raffield, L. M.; Czeisler, C. A.; Duffy, J. F.; Mignot, E.
Show abstract
Plasma proteomics is increasingly used for biomarker discovery and predictive modeling, yet diurnal protein trajectories remain insufficiently characterized. In our review of recent proteomic biomarker studies, 43% of the identified biomarkers had previously been reported to display 24-h rhythmicity. We demonstrate that ignoring these short-term dynamic effects compromises the robustness of reported models predicting health outcomes. We integrated a population-scale multi-ethnic longitudinal cohort with repeated measures over 10 years, with two cohorts of healthy adults undergoing frequent plasma sampling across days under controlled circadian, sleep and food-intake conditions. This design enabled estimation of short-term intraindividual variability (ST), long-term intraindividual variability (LT), population-level variability (POP) and genetic effects (GEN) across 7,289 protein targets. ST, LT, POP, and GEN define diverse protein trajectories, including rapid dynamics, long-term change, and individual-specific signatures. Using and generalizing this framework will facilitate covariate selection, study design, biomarker prioritization, and variability-aware modeling by users of proteomic data. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=92 SRC="FIGDIR/small/741635v1_ufig1.gif" ALT="Figure 1"> View larger version (26K): org.highwire.dtl.DTLVardef@55520eorg.highwire.dtl.DTLVardef@17e4bacorg.highwire.dtl.DTLVardef@9a0d0borg.highwire.dtl.DTLVardef@1ce88fb_HPS_FORMAT_FIGEXP M_FIG C_FIG
Obermiller, S. A.; Lipton, M. S.; Piehowski, P. D.; Bilbao, A.; McCue, L. A.; Prozapas, V. N.; Attah, I. K.
Show abstract
The functional complexity inherent in microbiomes complicates analytical approaches aimed at defining phenotype. As proteins are the functional effectors of microbiome phenotypes, improving the performance of mass spectrometry-based metaproteomics is critical to achieving the functional characterization of these systems. Data-independent acquisition (DIA) improves protein coverage and reduces data missingness when compared to data-dependent acquisition (DDA) in metaproteomics. However, the application of DIA to complex microbial systems remains constrained by analytical throughput and computational scalability. Here, we optimized LC-MS/MS acquisition parameters for both DDA and DIA using a model microbiome, demonstrating how DIA enables increased sample throughput without compromising quantitative performance. In addition, we demonstrated a computationally efficient, library-free DIA workflow that overcomes reliance on empirical spectral libraries. Our analytical and computational innovations establish a scalable and cost-effective pipeline for metaproteomics of complex microbial communities.